Papers by Daniel S. Weld
SpanBERT: Improving Pre-training by Representing and Predicting Spans (2020.tacl-1)
Copied to clipboard
| Challenge: | Pre-training methods like BERT mask individual words or subword units, but many tasks involve reasoning about relationships between two or more spans of text. |
| Approach: | They propose a pre-training method that masks contiguous random spans instead of random tokens to train the span boundary representations to predict the entire content of the masked span. |
| Outcome: | The proposed method outperforms BERT and its better-tuned baselines on span selection tasks and on coreference resolution tasks. |
VILA: Improving Structured Content Extraction from Scientific PDFs Using Visual Layout Groups (2022.tacl-1)
Copied to clipboard
| Challenge: | Recent work has improved extraction accuracy by incorporating elementary layout information, for example, each token’s 2D position on the page, into language model pretraining. |
| Approach: | They propose a method that explicitly models VIsual LAyout (VILA) groups, that is, text lines or text blocks, to further improve extraction accuracy. |
| Outcome: | The proposed methods show that inserting special tokens denoting layout group boundaries can lead to a 1.9% Macro F1 improvement in token classification. |